Skip to main content
eScholarship
Open Access Publications from the University of California

UCSF

UC San Francisco Electronic Theses and Dissertations bannerUCSF

Leveraging Artificial Intelligence for Bidirectional Data Harmonization and Developing Neighborhood Social Determinants of Health Archetype to Advance Chronic Disease Disparities Research

Abstract

The Multiple Chronic Disease Disparities Research (MCD-DR) is a National Institute on Minority Health and Health Disparities (NIMHD)-funded research consortium aimed at preventing and managing multiple chronic conditions among diverse populations affected by health disparities. Established through the Consolidated Appropriations Act of 2021, the Consortium includes 45 R01-level studies, primarily randomized controlled trials, from 11 P50 research centers, along with the University of California San Francisco (UCSF) Research Coordinating Center (RCC). A bidirectional data harmonization (DH) has been implemented to align survey data for secondary use across 41 studies, primarily clinical trials, thereby providing sufficient power to address the disproportionate burden of disease and health outcomes among these populations.The overall goal of this dissertation is to leverage artificial intelligence for bidirectional DH, and to develop neighborhood social determinants of health (SDOH) archetype, to advance chronic disease disparities research. The first chapter describes the Consortium DH pipelines, as well as challenges and opportunities encountered. The DH pipeline included developing common data elements (CDEs), creating working groups, considering data security and privacy, planning data transfer and management, and communicating with partners. Working with the MCD-DR Consortium highlights the need for a coordinated approach to balance comprehensiveness, manage variable heterogeneity, and validate measures across diverse groups. These efforts enhance research power and validity, providing a foundation for future initiatives to reduce health disparities and improve outcomes.The second chapter aims to develop a semi-automated DH pipeline using LLMs among eight studies from the eleven MCD-DR consortium P50 centers as a proof of concept. While LLM can help alleviate the burden of DH, optimizing the AI-implemented DH pipeline for bias, errors, cost and effort is critical. We utilized the Microsoft Azure OpenAI model (GPT-4o) hosted on the UCSF Versa platform. We first identified semantic clusters of each project’s questionnaire items that could be mapped to each RCC CDE semantic group based on the project’s and RCC’s CDE data dictionary. We compared the accuracy of matched pairs and CDE semantic groups with and without human-in-the-loop (HITL) during the semantic mapping. Subsequently, LLM was utilized to generate synthetic data for each project and coding for the variable transformation. We also utilized LLM to obtain advice on conducting variable mapping measured with different structures using the semantically mapped pairs across the projects. 90.1% of the items in the data dictionary were mapped to an average of 79 RCC CDE semantic groups per project. In the HITL semantic mapping, the accuracy for matched pairs ranged from 85.6% to 100.0% across studies with a mean 95.9% (mean CDE group accuracy 98.9%), whereas in the mapping without HITL, it ranged from 69.8% to 97.4% with a mean of 86.7% (mean CDE group accuracy 94.9%). This chapter demonstrated the potential of a human-interactive AI DH pipeline to augment researchers’ capabilities in expert areas during semantic mapping, as well as in generating coding systems for variable mapping and determining the final harmonization levels.The third chapter aims to identify neighborhood archetypes, patterns of characteristics, from diverse SDOH characteristics to inform the Consortium researchers and advance public health research. Using the most recent census tract-level dataset from the UCSF Health Atlas (2020 Census shape), neighborhood SDOH archetypes were developed using identification of domains with text embedding large language models, dimensional reduction with principal component analysis (PCA), and unsupervised autoencoders and Gaussian mixture models to identify archetypes in a cloud computing environment. For exploratory validation, the distribution of historical redlining, health outcome prevalence (15 conditions), and existing neighborhood indices, Powell-Wiley Neighborhood Deprivation Index and Child Opportunity Index, were compared according to identified neighborhood archetypes, while adjusting p-values for multiple testing. A total of 42 principal components across the five domains identified eleven neighborhood archetypes, which displayed heterogeneity in the distribution of the Neighborhood Deprivation Index and Child Opportunity Index subdomains in post hoc pairwise comparisons. Archetypes with higher deprivation had larger percentages of hazardous neighborhoods according to historical redlining compared to those with lower deprivation. Additionally, heterogeneous distributions of multiple chronic conditions among neighborhood archetypes suggested disproportionate burdens of health outcomes. In sum, this chapter demonstrated leveraging integrated data from the UCSF Health Atlas to identify and develop multifaceted neighborhood archetypes for sustainable deidentification and data safety.

Main Content

This item is under embargo until September 18, 2027.